Papers with NLP technology
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)
Copied to clipboard
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
Cross-lingual Few-Shot Learning on Unseen Languages (2022.aacl-main)
Copied to clipboard
| Challenge: | Large pre-trained language models have demonstrated the ability to obtain good performance on downstream tasks with limited examples in resource-rich languages. |
| Approach: | They propose to use a downstream sentiment analysis task to analyze the effectiveness of several few-shot learning strategies across 12 languages, including 8 unseen languages, to compare results. |
| Outcome: | The proposed model, XLM-R, gives the best performance on a task with few examples in resource-rich languages. |
Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender (2022.coling-1)
Copied to clipboard
| Challenge: | Current modeling of 3rd person pronouns ignores neopronoun phenomena like naive pronounes, which are not (yet) widely established. |
| Approach: | They propose to validate existing and novel approaches for modeling 3rd person pronouns in language technology and validate them through a survey. |
| Outcome: | The proposed model excludes non-binary users, while ignoring gender-specific phenomena. |
Evaluating the Diversity, Equity, and Inclusion of NLP Technology: A Case Study for Indian Languages (2023.findings-eacl)
Copied to clipboard
| Challenge: | In order for NLP technology to be widely applicable, fair, and useful, it needs to serve a diverse set of speakers across the world’s languages, be equitable, not unduly biased towards any particular language, and be inclusive of all users. |
| Approach: | They propose to use Gini coefficient to assess NLP across all three dimensions to assess diversity, equity, and inclusion across all languages. |
| Outcome: | The proposed evaluation paradigm assesses NLP technologies across all three dimensions and identifies the need for regional-specific choices in model building and dataset creation. |
Benchmarking the Simplification of Dutch Municipal Text (2024.lrec-main)
Copied to clipboard
| Challenge: | Text simplification (TS) is a technique that makes written information more accessible to all people, especially those with cognitive or language impairments. |
| Approach: | They propose to use English as a pivot language for simplification of Dutch medical and municipal texts. |
| Outcome: | The proposed approach improves on Dutch medical text, while the existing pipeline performs better on all metrics. |
NERetrieve: Dataset for Next Generation Named Entity Recognition and Retrieval (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a widely adopted NLP task . authors present three variants of NER task, with dataset to support them . |
| Approach: | They propose three variants of the NER task, together with a dataset to support them . they propose a move towards more fine-grained entities and zero-shot recognition . |
| Outcome: | The proposed model matches or surpasses existing models in NER tasks . the proposed model is based on a large, silver-annotated corpus of 4 million paragraphs . |
GlotLID: Language Identification for Low-Resource Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing web-mined datasets for low-resource languages have been useful for low resource NLP. |
| Approach: | They propose a model that identifies 1665 low-resource languages and a new model that is rigorously evaluated and reliable. |
| Outcome: | The proposed model outperforms baselines when balancing F1 and false positive rate (FPR). |
Impoverished Language Technology: The Lack of (Social) Class in NLP (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing work on socio-demographic factors has focused on how much a person's socioeconomic status affects their language production and perception. |
| Approach: | They propose to include socio-economic class in future natural language processing (NLP) research aimed at understanding relationships between socio-demographic factors and language production and perception. |
| Outcome: | The proposed definition of class can be operationalised by NLP researchers and argue for including socio-economic class in future language technologies. |